Abstract
Background: AI is increasingly used in clinical research, generating a growing need for robust critical appraisal tools to evaluate methodological quality, reporting standards, and potential biases. While traditional instruments exist for conventional clinical studies, specific tools designed for AI-based research are still emerging.
Objective: This dataset accompanies a scoping review that aimed to identify and describe existing critical appraisal tools, reporting frameworks, and bias classification systems applicable to clinical studies using AI, including chatbot-based interventions.
Methods: We systematically searched MEDLINE, Embase, CINAHL, PsycINFO, and IEEE Xplore from inception to April 2024. Eligible studies included those proposing or using tools for critical appraisal, reporting, quality assessment, or risk of bias in AI-related clinical research. Screening and extraction followed JBI and PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) recommendations. Data extraction combined human review with a supervised GPT-4o–based retrieval-augmented generation (RAG) process to enhance transparency and reproducibility. All AI-assisted outputs were verified independently by two reviewers.
Results: Seventy records were included: 46 reporting guidelines (comprising 26 guides for reporting AI studies, 16 critical appraisal tools, 2 quality assessment instruments, and 2 risk-of-bias tools), 9 bias classification or bias mitigation studies, and 15 chatbot evaluation studies. All datasets, extraction templates, and RAG prompts are publicly available to facilitate validation and reuse.
Conclusions: This dataset provides a comprehensive overview of critical appraisal and reporting tools for AI-based clinical research. It may support the development of standardized evaluation frameworks and promote transparency in future AI-assisted health studies.
Trial Registration: OSF Registries 10.17605/OSF.IO/ETYDS; https://osf.io/etyds/overview
doi:10.2196/85688
Keywords
Introduction
The use of predictive or generative AI in health research is rapidly growing [-]. As in traditional clinical studies, the methods used in AI-assisted studies can introduce systematic errors. The translation of AI-assisted evidence into clinical practice and research requires critical appraisal tools for clinical decision-makers and researchers [-]. We carried out a scoping review [] to identify existing tools for the critical appraisal of clinical studies that use AI and to examine the concepts and domains these tools explore. Our research question is framed using the PCC (population, concept, and context) framework [,]: the population includes clinical studies using AI; the concept refers to tools for critical appraisal and associated constructs such as quality, reporting, validity, risk of bias, and applicability; and the context is clinical practice. In addition, bias classification and chatbot assessment studies were included.
A total of 70 records were included in the review, comprising a heterogeneous group. Of these, 46 were tools (26 guides for reporting AI studies, 16 tools for critical appraisal, 2 tools for study quality, and 2 tools for risk of bias), 9 were papers focused on bias classification or mitigation, and 15 were chatbot studies (6 chatbot assessment studies and 9 systematic reviews of chatbot studies).
Just as traditional health research may include systematic errors that lead to biased or nongeneralizable results, so AI methods can introduce their own systematic errors at the design, data collection, training, or evaluation stages, which threaten the validity and reliability of AI models’ data analyses, findings, and conclusions. Such errors can arise from a number of different sources, including, but not limited to, flawed data, biased algorithms, and incorrect training. Health care decision-makers therefore need to be able to critically appraise AI studies to detect these problems so they can assess the certainty and relevance of the presented evidence.
Access to this dataset will help systematize the available tools and allow other researchers to compare and select appropriate ones. The dataset will also be of use for teaching purposes and risk-of-bias evaluation. In addition, we believe that it is good scientific practice to share data.
Our objective is to make the primary data used in our scoping review publicly available.
Methods
Overview
We searched medical and engineering databases (MEDLINE, Embase, CINAHL, PsycINFO, and IEEE) from inception to April 2024. We included primary clinical research that used tools for critical appraisal. Classic reviews and systematic reviews were included in the first phase of screening and used to identify new tools by forward snowballing. They were excluded in the second phase. We excluded nonhuman, computer, and mathematical research and letters, opinion papers, and editorials. We used Rayyan for screening [].
The protocol was previously registered with OSF []. We adhered to the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews; ) [] and the PRISMA-S (Preferred Reporting Items for Systematic Reviews Literature Search Extension) for reporting literature in systematic reviews [].
Data extraction for tools and bias was performed according to JBI recommendations [], and each study was rated by two observers. Discrepancies were resolved by discussion and consensus in an iterative process. (Data extraction tables and guidance are shown in .)
Data extraction for chatbot studies used a hybrid approach, combining the active involvement of a researcher with a retrieval-augmented generation (RAG) approach using custom ChatGPT instances based on the GPT-4o model, accessed directly through the ChatGPT web interface, with default model parameters (the RAG prompt template in Spanish is available in Spanish in and in English in ). Each PDF was uploaded and analyzed as an independent prompt, and all extracted information was subsequently reviewed by the primary investigator and independently validated by a second investigator to minimize hallucinations and ensure that the large language model was used solely as a facilitative extraction tool. No sensitive data were exposed. To promote transparency and reproducibility, the exact prompts used in the RAG process are incorporated into this dataset.
As an example to explain the process, the systematic review approach with RAG (in Spanish) was used to extract data from Oh et al []. A PDF was uploaded to ChatGPT (GPT-4o), and the output was provided in a PDF file (). The result was incorporated as a column in the rev_sis extraction XLS file, which can be found in column H in the dataset []. This procedure was repeated for all systematic review articles and for all primary research articles (rev_sis extraction and primary studies extraction, respectively). Both XLS files were transposed from columns to rows (matrix transposition) to build the final XLS files (S reviews table draft and primary studies table draft, respectively). The results were translated from Spanish to English by the reviewers. Both table drafts were reviewed and their results compared with the actual papers by LR-R and afterward verified by JBCL. The final tables were synthesized as shown in the scoping review, where they appear as Tables 4 and 5 [].
We identified 4392 records in the selected databases and registries. After eliminating 470 duplicates, 3922 records were screened by title and abstract, and 3803 were excluded. The remaining 119 underwent full-text screening, and 59 were excluded. The reasons for exclusion were as follows: 50 studies were systematic reviews, 7 met the exclusion criteria, and 2 did not meet the inclusion criteria. Full details are available in . Of the 50 systematic reviews, 42 used specific AI tools to assess the quality of the reviewed studies, and the tools retrieved were incorporated into the “records identified via other methods” category.
Twelve studies were identified in the EQUATOR Network library, and 4 additional studies were obtained from experts and organizations; therefore, 58 records were identified by other methods. Forty-eight of these were already captured in the 60 included studies from the search of electronic databases, leaving 10 additional studies to be included.
Ethical Considerations
Institutional review board approval was not applicable to this study and dataset. The scoping review built on this dataset has been accepted for publication in the Journal of Medical Internet Research.
Results
We identified 4392 records in databases and registries. After the selection process and inclusion of additional studies described in the Methods section, a total of 70 studies were included in the review (see scoping review []; ). Forty-six records reported tools (26 guidelines, 16 tools for critical appraisal, 2 tools for study quality, and 2 tools for risk of bias), and 9 were focused on bias classification or mitigation (). In addition, 15 records were chatbot studies, comprising 6 chatbot assessment studies and 9 systematic reviews of chatbot studies ().

Discussion
Overview
As explained in our main paper [], we conducted a comprehensive scoping review and identified 70 papers corresponding to the 3 proposed areas of research: tools for critical appraisal, bias and bias mitigation, and chatbot assessment studies. Although critical appraisal tools are the main focus of the review, types of AI bias were also included in the review because the validity (or absence of bias) is an important component of critical appraisal. Chatbot studies were included in the review because they represent an important recent, disruptive technology. The three areas together map the current landscape of evidence in the critical appraisal of clinical AI studies.
We selected critical appraisal as the main domain for the review because it is a wider and more inclusive concept than risk of bias, quality, or reporting, and it is more related to clinical practice. This decision required a change to the published protocol and was made after discussion.
Reporting guidelines are essential for authors in writing studies and for editors in maintaining consistency across publications. Critical appraisal tools are more focused on making judgments about the validity and applicability of evidence, and they are mainly intended for dissemination and teaching purposes. A paper may be of no use in a clinical setting, even if it is perfectly reported and its data are valid. Finally, both quality and risk of bias are precise concepts, and their tools are complex and designed as far as possible to avoid inconsistencies. Such tools are more suitable for use in research syntheses. Nevertheless, reporting, critical appraisal, risk of bias, and quality form a cluster of closely related constructs with overlapping areas.
Adequate reporting varies by the structure and type of study and is not only an editorial requirement but a part of study quality. Obviously, good reporting is a precondition to assess study quality, but there is also empirical evidence that some reporting flaws (and some nonreporting flaws) are associated with bias in the effect estimation [,]. Therefore, exploring reporting is essential to judge the validity of any study, as it facilitates study replication, risk-of-bias or quality assessments, interpretation of the results, and judgment of the value and applicability of the results in real clinical settings for individualized or collective decisions. It is also needed for the inclusion and assessment of studies in systematic reviews and for the evaluation of systematic reviews. Therefore, it is part of the critical appraisal process [,].
The overlap of reporting and critical appraisal was a source of inconsistency between raters when classifying papers in this scoping review. Iterative discussions were necessary to reach consensus. The most important criteria we used to classify papers within the critical appraisal category were the relevance of the question in the clinical context and a clear intent to help with applicability.
On the other hand, chatbot assessment studies are heterogeneous and inconsistent in their design, analysis, and reporting, so we used ChatGPT (GPT-4o) for data extraction; however, all outputs were independently reviewed by two authors against the original articles, and no major corrections to the extracted information were required. Therefore, we believe this was a valid and consistent procedure for data extraction. Nevertheless, we include the RAG and instructions for matrix transposition to enable other researchers to replicate the process.
A recent systematic review [] synthesized reporting guidelines as well as tools for basic and laboratory research. However, the search was conducted only through 2022. The available reporting guidelines should be harmonized, and the review would benefit from being updated or reformulated from a clinical standpoint.
There are some limitations of this scoping review. First, we used a general search that included all of our study’s questions; it was not specifically designed to search for bias and bias mitigation or for chatbot assessment. However, the absence of MeSH for chatbot studies and the heterogeneity of objectives, research questions, study design, devices, and analyses make searches for this type of study difficult. In addition, a potential limitation lies in the methods used to organize data extraction, as the application of large language models in evidence synthesis is novel, and formal standards for their integration are still under development. Finally, this field is evolving very quickly, so many conclusions drawn from existing evidence have a limited period of validity.
Critical appraisal tools are enormously varied, with different nuances and approaches, so selecting one can be very challenging. We believe that this topic deserves a qualitative synthesis to clarify the key elements for choosing the appropriate tool.
New risk-of-bias tools for AI in prognosis and diagnosis (such as QUADAS-AI [Quality Assessment of Diagnostic Accuracy Studies Using AI] and PROBAST+AI [Prediction Model Risk of Bias Assessment Tool]) and the PRISMA-AI (Preferred Reporting Items for Systematic Reviews and Meta-Analyses—Artificial Intelligence Extension) for systematic reviews are expected to be published, as is CHART (Chatbot Assessment Reporting Tool), a tool for reporting chatbot assessment studies. The AI extensions of other classic tools, such as the Cochrane risk-of-bias tool and ROBINS-I (Risk of Bias in Non-Randomized Studies of Interventions), among others, should be considered. On the other hand, the development of standards for the design, reporting, and assessment of chatbot assessment studies and chatbot health-advising studies is a clear gap in our toolbox and needs to be addressed.
In clinical practice, it is important to clarify the appropriate selection of tools for critical appraisal; furthermore, it is essential to develop teaching strategies to promote skills for the critical appraisal of AI-assisted studies, including understanding the types of bias to be tackled.
Conclusions
This dataset may be useful for other researchers who want to corroborate or further develop our results about critical appraisal tools for clinical studies that use AI.
Acknowledgments
Although we used ChatGPT (GPT-4o) for data extraction, we did not use generative AI for manuscript production.
Funding
No external financial support or grants were received from any public, commercial, or not-for-profit entities for the research, authorship, or publication of this article.
Conflicts of Interest
None declared.
Multimedia Appendix 2
ChatGPT retrieval-augmented generation prompt template (Spanish).
DOCX File, 17 KBMultimedia Appendix 3
ChatGPT retrieval-augmented generation prompt template (English).
DOCX File, 19 KBMultimedia Appendix 6
Included studies reporting critical appraisal tools and bias mitigation.
XLSX File, 159 KBChecklist 1 PRISMA-ScR checklist
DOCX File, 111 KBReferences
- Kaul V, Enslin S, Gross SA. History of artificial intelligence in medicine. Gastrointest Endosc. Oct 2020;92(4):807-812. [CrossRef] [Medline]
- Kohane IS. Injecting artificial intelligence into medicine. NEJM AI. Jan 2024;1(1). [CrossRef]
- Jayakumar S, Sounderajah V, Normahani P, et al. Quality assessment standards in artificial intelligence diagnostic accuracy systematic reviews: a meta-research study. NPJ Digit Med. Jan 27, 2022;5(1):11. [CrossRef] [Medline]
- Quirk J, Mac Donnchadha C, Vaantaja J, et al. Future implications of artificial intelligence in lung cancer screening: a systematic review. BJR Open. Oct 15, 2024;6(1):tzae035. [CrossRef] [Medline]
- Fleuren LM, Klausch TLT, Zwager CL, et al. Machine learning for the prediction of sepsis: a systematic review and meta-analysis of diagnostic test accuracy. Intensive Care Med. Mar 2020;46(3):383-400. [CrossRef] [Medline]
- Chakraborty C, Pal S, Bhattacharya M, Dash S, Lee SS. Overview of chatbots with special emphasis on artificial intelligence-enabled ChatGPT in medical science. Front Artif Intell. 2023;6:1237704. [CrossRef] [Medline]
- Barker TH, Stone JC, Sears K, et al. Revising the JBI quantitative critical appraisal tools to improve their applicability: an overview of methods and the development process. JBI Evid Synth. Mar 1, 2023;21(3):478-493. [CrossRef] [Medline]
- Moher D. Reporting guidelines: doing better for readers. BMC Med. Dec 14, 2018;16(1):233. [CrossRef] [Medline]
- Ibrahim H, Liu X, Rivera SC, et al. Reporting guidelines for clinical trials of artificial intelligence interventions: the SPIRIT-AI and CONSORT-AI guidelines. Trials. Jan 6, 2021;22(1):11. [CrossRef] [Medline]
- Crossnohere NL, Elsaid M, Paskett J, Bose-Brill S, Bridges JFP. Guidelines for artificial intelligence in medicine: literature review and content analysis of frameworks. J Med Internet Res. Aug 25, 2022;24(8):e36823. [CrossRef] [Medline]
- Cabello JB, Ruiz Garcia V, Torralba M, et al. Critical appraisal tools for evaluating artificial intelligence in clinical studies: scoping review. J Med Internet Res. Dec 8, 2025;27:e77110. [CrossRef] [Medline]
- Levac D, Colquhoun H, O’Brien KK. Scoping studies: advancing the methodology. Implement Sci. Sep 20, 2010;5(1):69. [CrossRef] [Medline]
- Aromataris E, Lockwood C, Porritt K, Pilla B, Jordan Z, editors. JBI Manual for Evidence Synthesis. JBI; 2024. [CrossRef]
- Ouzzani M, Hammady H, Fedorowicz Z, Elmagarmid A. Rayyan—a web and mobile app for systematic reviews. Syst Rev. Dec 5, 2016;5(1):210. [CrossRef] [Medline]
- Critical appraisal tool for artificial intelligence clinical studies. a scoping review. Open Science Framework. Apr 18, 2024. URL: https://osf.io/etyds/overview [Accessed 2026-09-07]
- Tricco AC, Lillie E, Zarin W, et al. PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation. Ann Intern Med. Oct 2, 2018;169(7):467-473. [CrossRef] [Medline]
- Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Syst Rev. Jan 26, 2021;10(1):39. [CrossRef] [Medline]
- Pollock D, Peters MDJ, Khalil H, et al. Recommendations for the extraction, analysis, and presentation of results in scoping reviews. JBI Evid Synth. Mar 1, 2023;21(3):520-532. [CrossRef] [Medline]
- Oh YJ, Zhang J, Fang ML, Fukuoka Y. A systematic review of artificial intelligence chatbots for promoting physical activity, healthy diet, and weight loss. Int J Behav Nutr Phys Act. Dec 11, 2021;18(1):160. [CrossRef] [Medline]
- Dechartres A, Trinquart L, Faber T, Ravaud P. Empirical evaluation of which trial characteristics are associated with treatment effect estimates. J Clin Epidemiol. Sep 2016;77:24-37. [CrossRef] [Medline]
- Dwan K, Altman DG, Clarke M, et al. Evidence for the selective reporting of analyses and discrepancies in clinical trials: a systematic review of cohort studies of clinical trials. PLoS Med. Jun 2014;11(6):e1001666. [CrossRef] [Medline]
- Simera I, Moher D, Hirst A, Hoey J, Schulz KF, Altman DG. Transparent and accurate reporting increases reliability, utility, and impact of your research: reporting guidelines and the EQUATOR Network. BMC Med. Apr 26, 2010;8(1):24. [CrossRef] [Medline]
- Kolbinger FR, Veldhuizen GP, Zhu J, Truhn D, Kather JN. Reporting guidelines in medical artificial intelligence: a systematic review and meta-analysis. Commun Med (Lond). Apr 11, 2024;4(1):71. [CrossRef] [Medline]
Abbreviations
| CHART: Chatbot Assessment Reporting Tool |
| PCC: population, concept, and context |
| PRISMA-AI: Preferred Reporting Items for Systematic Reviews and Meta-Analyses—Artificial Intelligence Extension |
| PRISMA-S: Preferred Reporting Items for Systematic Reviews and Meta-Analyses Literature Search Extension |
| PRISMA-ScR: Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews |
| PROBAST+AI: Prediction Model Risk of Bias Assessment Tool |
| QUADAS-AI: Quality Assessment of Diagnostic Accuracy Studies Using AI |
| RAG: retrieval-augmented generation |
| ROBINS-I: Risk of Bias in Non-Randomized Studies of Interventions |
Edited by Amaryllis Mavragani; submitted 13.Oct.2025; peer-reviewed by Andrea Conti, Seongsoon Kim; final revised version received 27.Apr.2026; accepted 03.Jun.2026; published 28.Aug.2026.
Copyright© Juan Bautista Cabello López, Vicente Ruiz García, Miguel Torralba, Miguel Maldonado Fernandez, María del Mar Úbeda-Carrillo, Eukene Ansuategi, Luis Ramos-Ruperto, José Ignacio Emparanza, Iratxe Urreta-Barallobre, María-Teresa Iglesias Gaspar, José Ignacio Pijoan Zubizarreta, Amanda J Burls. Originally published in JMIR Data (https://data.jmir.org), 28.Aug.2026.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Data, is properly cited. The complete bibliographic information, a link to the original publication on https://data.jmir.org/, as well as this copyright and license information must be included.

